Papers with internal mechanisms of bias exploitation
Debiasing Reward Models via Causally Motivated Inference-Time Intervention (2026.acl-long)
Copied to clipboard
| Challenge: | Existing approaches for mitigating spurious features in RMs focus on response length . Existing methods focus on RM activation, resulting in performance trade-offs . |
| Approach: | They propose a method that uses neurons to suppress spurious features in RMs at inference time. |
| Outcome: | The proposed method reduces sensitivity to spurious features without inducing performance trade-offs on RM benchmarks. |